Papers with direct model editing

2 papers
DELMAN: Dynamic Defense Against Large Language Model Jailbreaking with Model Editing (2025.findings-acl)

Copied to clipboard

Challenge: Existing safety mechanisms for Large Language Models (LLMs) are inadequate to protect against jailbreak attacks, resulting in performance degradation on general tasks.
Approach: They propose a method that directly updates a minimal set of relevant parameters to neutralize harmful behaviors while preserving the model’s utility.
Outcome: The proposed model outperforms baseline methods in mitigating jailbreak attacks while preserving the model’s utility.
Emptying the Ocean with a Spoon: Should We Edit Models? (2023.findings-emnlp)

Copied to clipboard

Challenge: a recent study has questioned the use of direct model editing for factual corrections in LLMs. aaron s. de stefano, a sociologist, says that model editing is not a systematic remedy for factuality.
Approach: They argue that direct model editing cannot be trusted as a remedy for LLM disadvantages . authors call for cautious promotion and application of model editing as part of LLM deployment process .
Outcome: The proposed method is not trusted as a remedy for the disadvantages inherent to LLMs, the authors argue . they argue that it opens risks by reinforcing the notion that models can be trusted for factuality .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations